feat(launch): re-add llm-d multinode support to infx.launch - #3613
Conversation
Add an llm-d driver that submits the job through the llm-d wrapper, attaches to the Slurm job, streams its log, checks its final state and stages results, agentic and eval artifacts. Route llmd-vllm multinode requests to it on gb200-nv and resolve DeepSeek-V4-Pro to the node-local numa1 checkpoint. Restore the GB200 DeepSeek-V4-Pro disagg wrapper. job.slurm no longer scancels its own allocation when the coordinator finishes; it stops the srun step and exits 0, so the job ends COMPLETED instead of CANCELLED, which infx.launch reports as a failure. Driver port by Ilya Markov from #2719.
|
Thanks for the contribution!
中文感谢你的贡献!
|
…est point Drop the per-model llm-d wrapper script. The driver passes topology, sequence lengths and concurrency to submit.sh and sets GPUS_PER_NODE, TIME_LIMIT, CONTAINER_IMAGE and worker counts itself. llmd-vllm supports only P/D disaggregated points. Re-add the dsv4-fp4-gb200-llmd-vllm low-latency 8k1k point (1P DEP8 + 1D TP8, conc 1) to exercise the path end to end.
Route every multinode llmd-vllm request to the llm-d driver. The driver requires slurm.squash and a staged checkpoint, which is what the whitelist stood in for.
/mnt/numa1 exists only on compute nodes, so the runner host cannot see the checkpoint. job.slurm already checks it on every allocated node.
…cher # Conflicts: # inferencex-e2e/infx/launch/drivers/__init__.py # inferencex-e2e/infx/launch/policy.py # inferencex-e2e/infx/launch/request.py # inferencex-e2e/perf-changelog.yaml
…cher # Conflicts: # inferencex-e2e/perf-changelog.yaml
The point validated the llm-d path end to end in e2e run 36906427804; it is not a benchmark to publish.
|
Claude finished @adibarra's task in 56s —— View job Review: PR #3613LGTM - no blocking issues found One note, not blocking: the PR description says the path is limited to gb200-nv through |
Add aggregate and disaggregated AgentX recipes on the Python llm-d launcher from #3613. Preserve explicit checkpoint selection, worker topology, offload metadata, and per-rank metrics collection. Keep successful Slurm completion distinct from failed or premature exits. Co-authored-by: Cam Quilici <[email protected]>
There was a problem hiding this comment.
Beyond the inline finding, I checked two other things the PR description's claims hinge on: policy.launch_path (inferencex-e2e/infx/launch/policy.py:80) actually routes any multinode llmd-vllm request to LaunchPath.LLMD regardless of cluster, not just gb200-nv — but srt/models.py only defines an llmd-vllm checkpoint entry for gb200-nv, so other clusters fail safely with a clear "stages no checkpoint" error rather than misbehaving. I also checked that drivers/llmd.py stages the image via backend.prepare_image() without calling run_container() — that's intentional, since submit.sh/job.slurm consume the squash reference directly through srun --container-image/pyxis rather than the launcher's normal container-run path.
Extended reasoning...
This PR re-adds llm-d multinode Slurm orchestration to the Python launcher (policy routing, a new driver, model/checkpoint resolution, and the job.slurm/server.sh scripts it shells out to) with no auth/crypto surface but real job-success/failure semantics at stake. A confirmed inline finding shows the job.slurm fix for the CANCELLED-on-success bug can itself mask a real failure as success when the decode coordinator exits early, because server.sh's EXIT trap touches the done marker unconditionally. I additionally verified two PR-description claims that looked suspicious in isolation (cluster gating in policy.py, and image staging without run_container in llmd.py) and found both fail safe rather than being functional bugs.
| run_until_bench_done() { | ||
| "$@" & | ||
| local step=$! rc=0 | ||
| while kill -0 "$step" 2>/dev/null; do | ||
| if [[ -f "$BENCH_DONE_MARKER" ]]; then | ||
| echo "[${SLURM_JOB_ID}] coordinator finished; stopping the srun step" | ||
| kill -TERM "$step" 2>/dev/null || true | ||
| break | ||
| fi | ||
| sleep 5 | ||
| done | ||
| wait "$step" || rc=$? | ||
| [[ -f "$BENCH_DONE_MARKER" ]] && return 0 | ||
| echo "[${SLURM_JOB_ID}] srun step exited rc=$rc before the coordinator finished" >&2 | ||
| return "$rc" |
There was a problem hiding this comment.
🔴 Any startup failure in the decode coordinator (EPP/Envoy not ready, prefill health timeout) now makes the whole llm-d job report success instead of failure, with no benchmark results. server.sh:303 has trap 'touch "$BENCH_DONE_MARKER"' EXIT, which fires on every exit of that branch, including the exit 1 failure paths, writing the same marker file as the real success touch at server.sh:573. job.slurm:82 [[ -f "$BENCH_DONE_MARKER" ]] && return 0 treats marker-exists as unconditional success, ignoring rc. Fix: use a distinct marker/exit path for genuine completion so job.slurm only succeeds when the benchmark actually ran, not merely when the coordinator exited for any reason.
Why this was flagged
Trigger: the decode-leader branch in server.sh (ROLE==decode && LWS_WORKER_INDEX==0) hits any of its exit 1 checks - EPP not binding within 60s (server.sh:384/388), Envoy /ready failing within 120s (server.sh:407/412), or prefill /health timing out within 300s (server.sh:509) - reached via job.slurm's run_until_bench_done -> srun -> server.sh. The EXIT trap at server.sh:303 touches $BENCH_DONE_MARKER on that exit, same file as the genuine-success touch at server.sh:573. job.slurm's run_until_bench_done (job.slurm:70-85) only checks file existence at line 82, not the step's rc, so it returns 0. set -eo pipefail (job.slurm:10) then lets the script finish normally, Slurm records COMPLETED, and drivers/llmd.py:137-138 computes rc=0 from that state; the result-JSON glob (llmd.py:140) finds nothing but the loop simply doesn't run, so no error is raised. On the base branch this same failure ends CANCELLED, which the launcher already treats as failure.
Verification: server.sh:303's trap 'touch "$BENCH_DONE_MARKER"' EXIT fires on every exit of the decode-coordinator branch, including the startup exit 1 failures at server.sh:384, 388, 407, 412, and 509, writing the same marker as the genuine-success touch at server.sh:573. job.slurm's [[ -f "$BENCH_DONE_MARKER" ]] && return 0 then discards the failing rc whenever that marker exists.
Add aggregate and disaggregated AgentX recipes on the Python llm-d launcher from #3613. Preserve explicit checkpoint selection, worker topology, offload metadata, and per-rank metrics collection. Keep successful Slurm completion distinct from failed or premature exits. Use golden synthetic acceptance for every AgentX MTP throughput role and launch accuracy evaluations separately with real verification. Co-authored-by: Cam Quilici <[email protected]>
#3613 routes multi-node llmd-vllm points to the llm-d driver instead of srt-slurm, so they name no srt-slurm recipe and need no CONFIG_FILE. The preflight now checks only the points the launcher hands to srt-slurm, and both read the same LLMD_FRAMEWORKS set from infx.launch.policy. #3613 让多节点的 llmd-vllm 点改由 llm-d driver 启动,而不是 srt-slurm, 因此这些点不引用 srt-slurm 配方,也不需要 CONFIG_FILE。预检现在只检查 launcher 交给 srt-slurm 的点,二者共用 infx.launch.policy 中的 LLMD_FRAMEWORKS。 Co-authored-by: Cursor <[email protected]>
Summary
Re-adds llm-d (
llmd-vllm) multinode support on the Python launcher from #3576, which dropped thelaunch_gb200-nv.shllm-d branch.llm-d keeps its own Slurm orchestration (
submit.sh→sbatch job.slurm→srun+ pyxis per node →server.sh). The launcher owns everything around it:policy.launch_path: multinodeframework == llmd-vllm→LaunchPath.LLMD, enabled ongb200-nvonly (LLMD_CLUSTERS); the table check requiresslurm.squash.drivers/llmd.py: resolves the checkpoint fromrunners.yaml, imports the image viabackend.prepare_image, runsllm-d/submit.shdirectly with the point's topology, sequence lengths and concurrency, attaches to the printed job id, streams the log, checks the final Slurm state, and stages result JSONs, agentic and eval artifacts plus the server-log bundle. No per-model wrapper script.srt/models.py:dsv4/fp4/llmd-vllmon gb200-nv →DeepSeek-V4-Pro@numa1, served asdeepseek-ai/DeepSeek-V4-Pro.Fix: successful jobs no longer end CANCELLED
job.slurmused toscancelits own allocation once the decode coordinator finished, so every successful run endedCANCELLED 0:0, which the launcher reports as a failure (run 36720223235 finished its benchmark and was still marked failed).job.slurmnow stops thesrunstep and exits 0, so the job endsCOMPLETED. A step that dies before the coordinator finishes still fails the job with its exit code.Driver port by @ilmarkov from #2719.
Testing
pytest infx/tests/launch infx/tests/clusters infx/tests/matrix infx/tests/workflows: 1230 passed; ruff cleanvalidate_perf_changelogagainstorigin/main: passesjob.slurmstep control: hung workers + done marker → exit 0; step failing first → its exit code